SemanticScuttle - klotz.me » Tags: machine learning+llm+deployment+vllm

Tags: machine learning* + llm* + deployment* + vllm*

0 bookmark(s) - Sort by: Date ↓ / Title /

From Flask to vLLM: How Model Inference has evolved (2017-2025)

The article discusses the evolution of model inference techniques from 2017 to a projected 2025, highlighting the progression from simple frameworks like Flask and FastAPI to more advanced solutions like Triton Inference Server and vLLM. It details the increasing demands on inference infrastructure driven by larger and more complex models, and the need for optimization in areas like throughput, latency, and cost.

2025-08-06 Tags: model inference, machine learning, deep learning, llm, vllm, triton, flask, fastapi, deployment by klotz

El Reg's essential guide to deploying LLMs in production

Running GenAI models is easy. Scaling them to thousands of users, not so much. This guide details avenues for scaling AI workloads from proofs of concept to production-ready deployments, covering API integration, on-prem deployment considerations, hardware requirements, and tools like vLLM and Nvidia NIMs.

2025-04-28 Tags: llm, ai, production engineering, inference engineering, deployment, vllm, nvidia, kubernetes, inference, api, scaling, gpu, machine learning by klotz

First / Previous / Next / Last / Page 1 of 0

SemanticScuttle - klotz.me

Tags: machine learning* + llm* + deployment* + vllm*

Linked Tags

Related Tags